Papers with task formulation

21 papers
TaxFree: a Visualization Tool for Candidate-free Taxonomy Enrichment (2022.aacl-demo)

Copied to clipboard

Challenge: In this paper, we present an open source system for taxonomy visualisation and automatic taxonomies enrichment without pre-defined candidates.
Approach: They propose an open source system for taxonomy visualisation and automatic taxonomie enrichment without pre-defined candidates on the example of WordNet-3.0.
Outcome: The proposed system can be used for visualisation and inspection of taxonomies without pre-defined candidates on WordNet-3.0.
DATE: Detecting Anomalies in Text via Self-Supervision of Transformers (2021.naacl-main)

Copied to clipboard

Challenge: Recent deep learning methods for anomalies in images learn better features of normality in an end-to-end self-supervised setting.
Approach: They propose to use a novel pretext task to learn a deep learning model for Anomaly Detection in text to train a model to discriminate between different transformations applied to visual data.
Outcome: The proposed method outperforms state-of-the-art methods on 20Newsgroups and AG News datasets in the semi-supervised setting and in the unsupervised setting.
Preventing Critical Scoring Errors in Short Answer Scoring with Confidence Estimation (2020.acl-srw)

Copied to clipboard

Challenge: Recent Short Answer Scoring systems use Quadratic Weighted Kappa (QWK) but it is unsatisfactory when measuring their effectiveness in actual usage.
Approach: They propose a task formulation of Short Answer Scoring (SAS) that matches actual usage and extracts as many scoring predictions that are not critical scoring errors (CSEs).
Outcome: The proposed system predicts scores with zero critical scoring errors (CSEs) for 50% of test data at maximum by filtering out low-reliability predictions on the basis of a certain confidence estimation.
Measuring the Effect of Influential Messages on Varying Personas (2023.acl-short)

Copied to clipboard

Challenge: a new task estimates the response a persona might have upon seeing a news message . a first benchmark dataset is used to evaluate the performance of the proposed task .
Approach: They propose a task to estimate the response a persona might have upon seeing a news message.
Outcome: The proposed task estimates the response a persona might have upon seeing a news message.
Narrative Question Answering with Cutting-Edge Open-Domain QA Techniques: A Comprehensive Study (2021.tacl-1)

Copied to clipboard

Challenge: Recent advances in open-domain question answering (ODQA) have led to human-level performance on many datasets.
Approach: They provide a comprehensive and quantitative analysis about the difficulty of book QA . they compare the results of their research with extensive ODQA experiments .
Outcome: The proposed model outperforms existing models on event-oriented questions on the NarrativeQA dataset.
S2SPMN: A Simple and Effective Framework for Response Generation with Relevant Information (D18-1)

Copied to clipboard

Challenge: Existing work on how to generate relevant and informative responses is focusing on how dialogue systems generate information from large dialogue corpus.
Approach: They propose to use dialogue corpus to generate relevant responses by using prototypes to extract semantic information from PMN.
Outcome: The proposed model outperforms classical and strong baseline models in generating relevant and informative responses.
Challenges in Information-Seeking QA: Unanswerable Questions and Paragraph Retrieval (2021.acl-long)

Copied to clipboard

Challenge: Existing pretrained language models have solved reading comprehension benchmarks, but datasets with information-seeking queries remain challenging.
Approach: They analyze why answering information-seeking queries is more challenging . they manually annotate 800 unanswerable examples across six languages .
Outcome: The proposed model outperforms human annotators on 800 unanswerable examples across six languages.
FIBER: Fill-in-the-Blanks as a Challenging Video Understanding Evaluation Framework (2022.acl-long)

Copied to clipboard

Challenge: Existing video understanding evaluation frameworks that use fill-in-the-blanks do not reflect real-world tasks.
Approach: They propose to use fill-in-the-blanks as a video understanding evaluation framework and introduce a novel dataset that collects multiple perspectives on the same video.
Outcome: The proposed framework does not share the weaknesses of the current state-of-the-art language-informed video understanding tasks, namely: (1) video question answering using multiple-choice questions, where models perform relatively well because they exploit linguistic biases in the task formulation; (2) video captioning, which relies on an open-ended evaluation framework that is often inaccurate because system answers may be perceived as incorrect if they differ in form from the ground truth.
OLMES: A Standard for Language Model Evaluations (2025.findings-naacl)

Copied to clipboard

Challenge: Existing models claim to perform better on tasks measuring model capabilities, but there is no standard setup for reproducible evaluations.
Approach: They propose a document that is documented and practical for reproducible LLM evaluations and includes recommendations from existing literature and new experiments.
Outcome: The proposed standard identifies and reviews the varying factors in evaluation practices adopted by the community, such as prompt formatting, choice of in-context examples, probability normalizations, and task formulation.
COLA: Contextualized Commonsense Causal Reasoning from the Causal Inference Perspective (2023.acl-long)

Copied to clipboard

Challenge: Existing efforts to detect commonsense causation from the causal inference perspective are inadequate to seize commonsensical causations.
Approach: They propose a task to detect commonsense causation between two events in context . they propose 'contextualized commons sense causal reasoning' framework that uses covariates to remove confounding effects .
Outcome: The proposed framework can detect commonsense causality more accurately than baselines.
SOCCER: An Information-Sparse Discourse State Tracking Collection in the Sports Commentary Domain (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for state tracking are limited and state changes are less densely distributed over utterances.
Approach: They propose to turn to simplified, fully observable systems that show some of these properties.
Outcome: The proposed system shows that state changes occur infrequently while messages are "chatter" it allows for rich descriptions of state while avoiding the complexities of other settings.
ForecastQA: A Question Answering Challenge for Event Forecasting with Temporal Text Data (2021.acl-long)

Copied to clipboard

Challenge: Existing automated forecasting studies rely on structured data to predict future events.
Approach: They propose a question-answering task that limits access to unstructured text data . they use a crowdsourced dataset to form a restricted-domain, multiple-choice, question-announcement task .
Outcome: The proposed model achieves 61.0% accuracy on the dataset, which still lags behind human performance by about 19%.
Cross-lingual Contextualized Phrase Retrieval (2024.findings-emnlp)

Copied to clipboard

Challenge: Phrase-level dense retrieval has shown many appealing characteristics in downstream NLP tasks.
Approach: They propose a task formulation of dense retrieval, cross-lingual contextualized phrase retrieval . they extract pairs of cross-linguistic phrases using word alignment information .
Outcome: The proposed task formulation surpasses baselines on the phrase retrieval task and a downstream task, i.e., machine translation, and achieves top-1 accuracy 13 points higher.
Adaptive Text Anonymization: Learning Privacy-Utility Trade-offs via Prompt Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for anonymizing textual documents lack flexibility to adapt to diverse requirements.
Approach: They propose a task formulation in which anonymization strategies are automatically adapted to specific privacy–utility requirements.
Outcome: The proposed framework achieves better privacy–utility trade-off than existing baselines on open-source language models while remaining computationally efficient and effective on larger closed-source models.
Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements (2020.emnlp-main)

Copied to clipboard

Challenge: Existing tools for examining and fixing missing captions are lacking in mobile UIs.
Approach: They propose a task for automatically generating language descriptions for UI elements from multimodal input including both the image and structural representations of user interfaces.
Outcome: The proposed task can generate captions from image and structural representations of UI elements.
A Dataset for Tracking Entities in Open Domain Procedural Text (2020.emnlp-main)

Copied to clipboard

Challenge: Existing tasks require only a small set of attributes to track state changes in procedural text.
Approach: They propose a task where given a procedural text as input, the task is to generate a set of state change tuples for each step.
Outcome: The proposed task generates state change tuples from a set of pre-defined attributes for each step and predicts them from an open vocabulary.
Come hither or go away? Recognising pre-electoral coalition signals in the news (2021.emnlp-main)

Copied to clipboard

Challenge: In this paper, we decompose the task of recognizing from the news coverage leading up to an election the (un)willingness of political parties to form a coalition into two related, but distinct tasks.
Approach: They propose a task of recognizing from news coverage the (un)willingness of political parties to form a coalition from text and a sub-task of predicting the polarity of the signal.
Outcome: The proposed approach improves over a strong monolingual transfer learning baseline.
Goal-Driven Explainable Clustering via Language Descriptions (2023.emnlp-main)

Copied to clipboard

Challenge: Existing formulations neither consider the users’ goals nor explain clusters’ meanings.
Approach: They propose a task formulation that represents both the goal and the explanations as free-form language descriptions.
Outcome: The proposed method produces more accurate and goal-related explanations than previous methods.
Don’t Just Say “I don’t know”! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies investigate ways to refuse to answer unknown questions . Large Language Models (LLMs) display a significant level of overconfidence when answering questions that they are aware of.
Approach: They propose a self-alignment method to utilize Large Language Models to enhance its response-ability to unknown questions.
Outcome: The proposed method is superior to baseline methods on four types of unknown questions.
Dual-Stage Multi-Task Syntax-Oriented Pre-Training for Syntactically Controlled Paraphrase Generation (2024.findings-acl)

Copied to clipboard

Challenge: Syntactically controlled paraphrase generation (SCPG) aims to generate sentences with syntactic structures resembling given exemplars.
Approach: They propose a dual-stage multi-task pre-training scheme that uses a series of structure-oriented and syntax-oriented tasks to generate sentences with syntactic structures resembling given exemplars.
Outcome: The proposed method outperforms existing methods on all possible variants of SCPG tasks and significantly outperformed the popular T5 model.
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Empathy plays a pivotal role in fostering prosocial behavior, often triggered by the sharing of personal experiences through narratives.
Approach: They propose to use contrastive learning with masked LMs and supervised fine-tuning with large language models to improve empathy understanding in NLP models.
Outcome: The proposed methods show that there is low agreement among annotators and that cultural differences are a factor in their interpretation of empathy.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations